Papers with base model

143 papers
Boosting Dialog Response Generation (P19-1)

Copied to clipboard

Challenge: Neural models generate the most common and generic responses all the time . Empirical results show that our method can significantly improve the diversity of responses generated by sequence-to-sequence models.
Approach: They propose an iterative training process and ensemble method based on boosting to improve the diversity of responses generated by neural models.
Outcome: Empirical results show that the proposed method significantly improves diversity and relevance of responses generated by all models.
A Systematic Analysis of Base Model Choice for Reward Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Reinforcement learning from human feedback (RLHF) and reward modeling are key to training powerful large language models (LLMs).
Approach: They propose to combine RLHF and reward modeling to boost model selection . they also demonstrate that a small set of benchmarks could be combined to boost the model selection.
Outcome: The results show that the model selection can be improved by up to 14% compared to the most common (default) choice.
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance (2024.eacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) have an array of reasoning capabilities but face limitations such as error propagation and hallucination.
Approach: They propose to use a LLAMA-2 13B CHAT model to act as a task router and task solver to offload certain reasoning steps to external tools that are more suited for the task.
Outcome: The proposed model improves by 35.2% and 5.06% over baseline models and strong GPT-3.5 results.
Constraining word alignments with posterior regularization for label transfer (2022.naacl-industry)

Copied to clipboard

Challenge: Unsupervised word alignments are not always possible in industrial NLP pipelines, where multilingual annotation guidelines are complex and deviate from semantic consistency due to various factors.
Approach: They propose to constrain word alignment models to remain consistent with both source and target annotation guidelines by leveraging posterior regularization and labeled examples.
Outcome: The proposed model improves on the multiATIS++ dataset over AWESoME, and even a small amount of target language annotations can help.
When Language Model Meets Private Library (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing language models have been pre-trained on large-scale code corpora and generate decent code snippets.
Approach: They propose a framework that can provide pre-trained language models with the ability to generate code using private libraries.
Outcome: The proposed framework can generate code using private libraries using off-the-shelf language models or pre-trained models on code corpus containing API information.
Is Micro Domain-Adaptive Pre-Training Effective for Real-World Operations? Multi-Step Evaluation Reveals Potential and Bottlenecks (2026.eacl-industry)

Copied to clipboard

Challenge: Domain-adaptive pre-training (DAPT) is one approach for enabling LLMs to handle unseen knowledge.
Approach: They propose to disentangle the answering process into three subtasks and evaluate the performance of each subtask.
Outcome: The proposed model resolves the elicitation task that the base model struggled with but does not resolve other subtasks.
Progressive Self-Training with Discriminator for Aspect Term Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to extract aspect terms from review sentences are limited due to lack of annotated data.
Approach: They propose to refine conventional self-training to progressive self-teaching to reduce noise . they use a discriminator to filter the noisy pseudo-labels.
Outcome: The proposed model outperforms baseline models and achieves state-of-the-art performance on four SemEval datasets.
Synthetic Propaganda Embeddings To Train A Linear Projection (D19-50)

Copied to clipboard

Challenge: Using contextualized token embeddings, we can extract features of propaganda from contextualized embeddnings without fine-tuning the large parameters of the base model.
Approach: They propose a method for detecting fine-grained categories of propaganda in text by generating synthetically generated embeddings from pre-trained language models.
Outcome: The proposed method is used in the first shared task in fine-grained propaganda detection at NLP4IF as Team Stalin.
Persona-driven Simulation of Voting Behavior in the European Parliament with Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models exhibit a progressive left-leaning bias, but can also produce behavior that aligns with socioeconomic groups.
Approach: They analyze whether persona prompting can accurately predict individual voting decisions . they find that they can simulate the voting behavior of European Parliament members reasonably well .
Outcome: The proposed model can predict the voting behavior of European Parliament members reasonably well, with a weighted F1 score of approximately 0.793.
Towards Efficient Dialogue Processing in the Emergency Response Domain (2023.acl-srw)

Copied to clipboard

Challenge: Adapters perform dialogue act classification and domain-specific slot tagging in the emergency response domain.
Approach: They propose to build a system that performs dialogue act classification and domain-specific slot tagging while being efficient, flexible and robust.
Outcome: The proposed model performs well in the emergency response domain while being efficient, flexible and robust.
A Stable and Effective Learning Strategy for Trainable Greedy Decoding (D18-1)

Copied to clipboard

Challenge: Existing methods for decoding text using beam search are expensive and require reinforcement learning.
Approach: They propose a method that allows us to reap the full benefits of beam search with no additional computational cost.
Outcome: The proposed method outperforms greedy decoding and beam search on machine translation tasks with minimal computational cost.
Exploring Fine-Tuning for In-Context Retrieval and Efficient KV-Caching in Long-Context Language Models (2026.eacl-short)

Copied to clipboard

Challenge: Long-Context Language Models (LCLMs) can encode entire document collections, offering a strong alternative to retrieval-augmented generation (RAG).
Approach: They propose to use LCLMs to encode documents with context windows of millions of tokens to improve their performance.
Outcome: The proposed training strategies improve long-context performance and their robustness under compression techniques.
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework (2024.emnlp-demo)

Copied to clipboard

Challenge: Text-to-image (T2I) diffusion models are popular for image manipulation, but also for video generation.
Approach: They propose a novel T2I diffusion model based on latent diffusion that extends the base model for various applications.
Outcome: The proposed model achieves high quality and photorealism and is 3 times faster than the base model.
TARo: Token-level Adaptive Routing for LLM Test-time Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance.
Approach: They propose to use token-level Adaptive Routing to steer frozen LLMs toward structured reasoning entirely at inference time.
Outcome: Extensive experiments show that TARo significantly improves reasoning performance by up to +22.4% over base model and +8.4% .
IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing (2024.emnlp-industry)

Copied to clipboard

Challenge: Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers.
Approach: They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc.
Outcome: The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination.
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing work adopts separate modules for retrieval and generation, which may be suboptimal since the retrieval task and generation task cannot benefit from each other to improve performance.
Approach: They propose a backbone-shared RAG framework that uses a domain-specific corpus to continuously pre-train a model and then trains two plug-and-play Low-Rank Adaptation modules based on the shared backbone to minimize retrieval and generation losses respectively.
Outcome: The proposed framework outperforms baseline models by 5% and 13% in Hit@3 upon two datasets in retrieval evaluation and by 23% in terms of BLEU-3 in generation evaluation.
PE-QAT: Parameter-Efficient Quantization-Aware Training for Large Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Quantization Aware Training (QAT) is expensive to train and unscalable to large models.
Approach: They propose a parameter-efficient framework targeting per-channel 4-bit weight-activation quantization of large language models.
Outcome: The proposed framework preserves accuracy within 0.11 percentage points of the full-precision baseline on Llama-2-7B zero-shot tasks while training only 1.26% of total parameters.
Entity Commonsense Representation for Neural Abstractive Summarization (N18-1)

Copied to clipboard

Challenge: Current ELS’s are not sufficiently effective, possibly introducing unresolved ambiguities and irrelevant entities.
Approach: They propose an off-the-shelf entity linking system to extract linked entities and propose Entity2Topic (E2T) module attachable to a sequence-to-sequence model that transforms a list of entities into a vector representation of the topic of the summary.
Outcome: The proposed model improves the performance of the Gigaword and CNN summarization datasets by at least 2 ROUGE points.
ProConSuL: Project Context for Code Summarization with LLMs (2024.emnlp-industry)

Copied to clipboard

Challenge: Experimental results show that ProConSuL significantly improves code summaries and reduces the number of hallucinations.
Approach: They propose a framework to provide a large language model with precise information about the code structure from program analysis methods.
Outcome: The proposed framework significantly improves code summaries and reduces hallucinations compared to the base model.
Syntactic and Semantic-driven Learning for Open Information Extraction (2020.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that our approach significantly outperforms the supervised counterparts, and can even achieve competitive performance to supervised state-of-the-art (SoA) model.
Approach: They propose a syntactic and semantic-driven learning approach that can learn open IE models without human-labelled data by leveraging syntakic and semantic knowledge as noisier, higher-level supervision.
Outcome: The proposed approach outperforms supervised counterparts and can achieve competitive performance to supervised state-of-the-art models.
Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers (2026.acl-demo)

Copied to clipboard

Challenge: Existing post-training pipelines that generate QA pairs require costly expert annotation and synthetic data that drops evidence structure.
Approach: They propose a system that converts raw biomedical papers into evidence-enriched training sets and a domain-specialized VLM.
Outcome: Ryze synthesizes QA pairs with complete supporting evidence, reduces layout and OCR errors . the system outperforms the base model on LAB-Bench and surpasses GPT-5.2 by +3.8%.
Domain Adaptation of Foundation LLMs for e-Commerce (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have greatly improved the performance on most natural language tasks, and often show surprisingly good zero-shot generalization to new domains.
Approach: They propose to continuously pretrain the Llama 3.1 base models on 1 trillion tokens of e-commerce data to introduce domain specific knowledge into the model while at the same time keeping the general capabilities intact.
Outcome: The proposed model can be adapted to the new domain without sacrificing performance on general domain tasks.
IndiFoodVQA: Advancing Visual Question Answering and Reasoning with a Knowledge-Infused Synthetic Data Generation Pipeline (2024.findings-eacl)

Copied to clipboard

Challenge: Large Vision Language Models lack domain-specific data for reasoning on complex problems.
Approach: They propose to use explicit knowledge-infused questions, answers, and reasons to answer and reason upon the questions.
Outcome: The proposed model improves by 25% over the baseline model.
Improving Factuality of Abstractive Summarization without Sacrificing Summary Quality (2023.acl-short)

Copied to clipboard

Challenge: Recent studies have shown that most abstractive summarization models are unfaithful and suffer from a wide range of hallucination.
Approach: They propose a candidate summary generation and ranking technique to improve summary factuality without sacrificing quality.
Outcome: The proposed method shows that the model trained using the proposed method improves on factuality and similarity-based metrics without conflicting with the model.
Where to start? Analyzing the potential value of intermediate models (2023.emnlp-main)

Copied to clipboard

Challenge: a finetuned model may be better base models than the vanilla pretrained model . this scheme, often referred to as intertraining, is the focus of the present work .
Approach: They propose a scheme to analyze the potential intertraining gain independently for the target dataset and for a base model being considered as a starting point.
Outcome: The proposed model is strong even if training data was not aligned with target dataset.
FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking (2026.acl-industry)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly being adopted across various domains, but adoption in the regulated banking industry is limited due to their tendency to hallucinate, exhibit over-agreeable behavior, and lack alignment with domain-specific knowledge and constraints.
Approach: They propose a framework for training grounded domain-specific LLMs that optimizes answer quality, citation grounding, and calibrated refusal under real-world deployment constraints.
Outcome: The proposed model outperforms GPT-4.1 on citation grounding and calibrated refusal under real-world deployment constraints.
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Large-scale language models (LLMs) like ChatGPT have demonstrated impressive abilities in generating responses based on human instructions. however, their use in the medical domain can be challenging due to their lack of specific, in-depth knowledge.
Approach: They propose a system that integrates authoritative medical textbooks into LLMs’ framework using plug-and-play modules.
Outcome: The proposed system outperforms the specialized Med-PaLM 2 model on three medical QA tasks by 11.6% to 16.6%.
FedLFC: Towards Efficient Federated Multilingual Modeling with LoRA-based Language Family Clustering (2024.findings-naacl)

Copied to clipboard

Challenge: Existing frameworks for multilingual modeling face communication costs and parameter interference conflicts.
Approach: They propose a communication-efficient federated learning framework with low-rank adaptation and language family clustering for Multilingual Modeling (MM) they maintain the weights of the base model, updating the lightweight Low-rank adapt parameters to minimize communication costs.
Outcome: The proposed model outperforms the baseline models in performance and reduces communication overhead.
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data (2024.emnlp-industry)

Copied to clipboard

Challenge: Large language models exhibit significant performance discrepancies between high- and low-resource languages.
Approach: They present an open-source multilingual LLM with 8 billion parameters and a multilingual instruction dataset.
Outcome: The proposed model achieves consistent multilingual representations across languages.
Learning Shortcut Models for Efficient Recursive Reasoning (2026.acl-srw)

Copied to clipboard

Challenge: Recent research shows that Transformer-style models can be made more efficient by sharing parameters over blocks.
Approach: They propose a framework for distilling latent reasoning into a multiscale jump model that enables flexible test-time compute.
Outcome: Experiments on ARC-AGI show that the proposed model achieves competitive accuracy compared to recursive baselines while requiring fewer sequential updates.
Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms (2026.acl-industry)

Copied to clipboard

Challenge: Existing research focuses on enhancing large language models through scaling laws or fine-tuning strategies, but ignores the potential of using agent paradigms to compensate for the inherent weaknesses of small models.
Approach: They propose to use structured agent frameworks to improve effectiveness over direct prompting . they also propose to employ routing-based multi-agent systems with collaborative capabilities .
Outcome: The proposed model significantly outperforms direct prompting with single-agent systems . the proposed model is more reliable and cost-effective than other models .
Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer Models (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in NLP have led to the use of pre-trained Transformer models for transfer learning tasks becoming the most common way to solve target tasks.
Approach: They propose a 3-phase technique to adjust a base model for a classification task by adapting the model’s signal to the data distribution and a new data augmentation approach for Supervised Contrastive Learning to correct the unbalanced datasets.
Outcome: The proposed method is compared with other methods and compares it with other approaches.
Improving Long-Tail Relation Extraction with Collaborating Relation-Augmented Attention (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to handle wrong labeling and long-tail relations are labor-intensive and scarce training data.
Approach: They propose a neural network to handle wrong labeling and long-tail relations by collaborating relation-augmented attention.
Outcome: The proposed neural network improves the state-of-the-art on the NYT dataset .
SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents (2026.findings-eacl)

Copied to clipboard

Challenge: Large language model (LLM) based search agents are more likely to produce harmful outputs than base models.
Approach: They propose a query-level shaping term that rewards safe queries and penalizes unsafe ones.
Outcome: The proposed approach reduces harmfulness by over 70% across three red-teaming datasets while producing safe, helpful responses.
SAND: Boosting LLM Agents with Self-Taught Action Deliberation (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Model (LLM) agents finetuned with supervised finetuning may over-commit towards seemingly plausible but suboptimal actions due to limited action space exploration.
Approach: They propose a self-taught actioN deliberation framework that allows LLM agents to explicitly deliberate over candidate actions before committing to one.
Outcome: The proposed framework outperforms state-of-the-art methods on two representative interactive agent tasks and achieves an average 20% improvement over initial finetuning.
Uncertainty-Aware Label Refinement for Sequence Labeling (2020.emnlp-main)

Copied to clipboard

Challenge: Conditional random fields (CRF) for label decoding have been a problem for many tasks.
Approach: They propose a two-stage label decoding framework that model long-term label dependencies while being much more computationally efficient.
Outcome: The proposed method outperforms the CRF-based methods and greatly accelerates the inference process.
Unfamiliar Finetuning Examples Control How Language Models Hallucinate (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) generate plausible-sounding responses that are factually incorrect.
Approach: They propose an approach to learn more reliable reward models by modifying how unfamiliar finetuning examples are supervised to influence model responses to unfamiliar queries.
Outcome: The proposed approach improves the efficacy of RL factuality finetuning in long-form biography and book/movie plot generation tasks.
Unsupervised Domain Adaptation Method with Semantic-Structural Alignment for Dependency Parsing (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for dependency parsing are often of the pseudo-annotation type, but they fail to consider the change of model structure for domain adaptation.
Approach: They propose a method that accomplishes unsupervised cross-domain dependency parsing without using labeled data.
Outcome: The proposed method achieves consistent performance improvement on CODT1 and CTB9 domains.
CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: coding tasks require generated code to be fully executable and functionally correct . current agentic approaches struggle with multi-stage planning, generating, and debugging .
Approach: They propose a framework for LLM agents to efficiently explore the search space in different stages of the code generation process.
Outcome: The proposed framework achieves top results on 7 code generation benchmarks and a 31.9% solving rate on the SWEBench benchmark.
High-quality argumentative information in low resources approaches improve counter-narrative generation (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that fine-tuning improves the performance of language models . large language models generate acceptable texts in a number of scenarios, a study shows .
Approach: They show that fine-tuning improves the task of hate speech counter-narrative generation . they provide a subset of arguments and a good base model is required for the fine-uning to have a positive impact.
Outcome: The proposed model produces counter-narratives that are as satisfactory as the whole set.
TokAlign: Efficient Vocabulary Adaptation via Token Alignment (2025.acl-long)

Copied to clipboard

Challenge: Tokenization is a foundational step for Large Language Models (LLMs) but low compression rate of vanilla tokenizers decelerates training and inference process.
Approach: They propose a method to replace the vocabulary of Large Language Models (LLMs) by learning a one-to-one mapping matrix for token IDs.
Outcome: The proposed method significantly improves multilingual text compression rates and vocabulary initialization for Large Language Models.
STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that supervised training is still necessary for complex reasoning tasks.
Approach: They propose a method to integrate uncertainty-based active learning and LoRA to effectively integrate the two methods.
Outcome: The proposed approach outperforms baseline models on three reasoning tasks.
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for domain-specific reasoning with large language models require updating parameter updates.
Approach: They propose a plug-and-play intervention framework that adaptively steers LLM reasoning in activation space.
Outcome: The proposed framework achieves zero-shot accuracy improvements of 3.4–6.5% over the base model while outperforming chain-of-thought-style reasoning with 2–3 higher token efficiency and robust accuracy gains.
Transferable Post-training via Inverse Value Learning (2025.naacl-long)

Copied to clipboard

Challenge: Existing algorithms for post-training large datasets are requiring a large computational effort.
Approach: They propose to model the changes at logits level during post-training using a separate neural network . they demonstrate that the value network can be seamlessly integrated with another pre-trained model .
Outcome: The proposed model can be integrated with another pre-trained model during inference, enabling similar capability enhancements.
Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based Games (2020.emnlp-main)

Copied to clipboard

Challenge: Reinforcement Learning methods for text-based games fail to generalize on unseen games, especially in small data regimes.
Approach: They propose a Context Relevant Episodic State Truncation method for irrelevant token removal in observation text for improved generalization.
Outcome: The proposed method shows that it can generalize on unseen games using 10x-20x fewer training games compared to previous state-of-the-art methods despite requiring fewer number of training episodes.
SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual Learning (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved remarkable success across diverse tasks through large-scale pretraining.
Approach: They propose a framework that filters noisy components from LoRA updates via subspace similarity with the base model.
Outcome: The proposed framework improves accuracy by 12%, reduces forgetting by 29%, and filters out over 30% of LoRA parameters identified as noisy.
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages (2026.acl-long)

Copied to clipboard

Challenge: Continued pretraining (CPT) is a practical route to language adaptation, but improvements on demanding capabilities such as mathematical reasoning are limited.
Approach: They propose to use CPT to adapt large language models to African languages . they use math, code, and synthetic translated data to analyze their models .
Outcome: The proposed models improve on multilingual benchmarks and document-level translation.
Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models (2022.coling-1)

Copied to clipboard

Challenge: Recent results show that the mix-of-experts architecture is parameter inefficient . large-scale pre-trained language models can achieve excellent performance in many NLP tasks.
Approach: They propose to build a parameter-efficient mix-of-experts architecture by sharing information across experts.
Outcome: The proposed architecture increases model capacity without increasing computation costs.
Cross-Lingual Summarization with Pseudo-Label Regularization (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to cross-lingual summarization use only a single reference, resulting in an underrepresented hypothesis space.
Approach: They propose to use pseudo-labels to regularize cross-lingual summarization training by combining a single reference and a network to perform the model training.
Outcome: The proposed approach significantly improves over gold reference training in XLS with 8 languages from different families.
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs (2026.findings-acl)

Copied to clipboard

Challenge: prevailing taxonomies neglect robustness and honesty, yielding safer-on-paper but less useful systems.
Approach: They propose a soft-gating pipeline where a guardian predicts a binary risk label plus a concise explanation and prepends this advice to the original query for re-inference.
Outcome: The proposed model maintains safety while reducing over-refusal.
SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization (2022.acl-long)

Copied to clipboard

Challenge: Sequence-to-sequence neural networks have enabled great progress in abstractive summarization.
Approach: They propose to train a second-stage model performing re-ranking on a set of summary candidates by using a mixture of experts.
Outcome: The proposed model outperforms the base model on CNN- DailyMail, XSum and Reddit TIFU with a base PEGASUS.
ELAINE-medLLM: Lightweight English Japanese Chinese Trilingual Large Language Model for Bio-medical Domain (2025.coling-main)

Copied to clipboard

Challenge: Existing bilingual or multilingual medical LLMs are limited in multilingual data and therefore perform poorly in non-English languages such as Japanese and Chinese.
Approach: They propose to use a trilingual (English, Japanese, Chinese) large language model adapted for the bio-medical domain to harness the knowledge and abilities of the base model.
Outcome: The proposed model can support English, Japanese, and Chinese and is adapted for a bio-medical domain.
Offline Preference Optimization via Maximum Marginal Likelihood Estimation (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to align Large Language Models with human preferences are complex and unstable.
Approach: They propose a new approach that maximizes the marginal log-likelihood of a preferred text output by using the preference pair as samples for approximation.
Outcome: The proposed approach maximizes the marginal log-likelihood of a preferred text output, using the preference pair as samples for approximation, and forgoes the need for both an explicit reward model and entropy maximization.
TableLlama: Towards Open Large Generalist Models for Tables (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for interpreting, augmenting, and querying semi-structured tables require pretraining on tables or special model architecture design.
Approach: They construct a dataset with a variety of tables and tasks for instruction tuning and evaluating LLMs.
Outcome: The proposed model achieves comparable or better performance on 7 out of 8 in-domain tasks compared with the base model on 6 out-of-domain datasets.
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation (2025.emnlp-main)

Copied to clipboard

Challenge: Extensive experiments on challenging mathematical reasoning benchmarks demonstrate that these human-inspired strategies synergistically and significantly enhance performance.
Approach: They propose to use Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation to improve model performance.
Outcome: Extensive experiments on mathematical reasoning benchmarks show that the proposed strategies synergistically and significantly improve performance over the baseline model.
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work focuses on domain-specific enhancements during fine-tuning, the challenge of which lies in catastrophic forgetting of knowledge across other domains.
Approach: They propose a data composition framework that allows LLMs to enhance their multi-domain capabilities during supervised fine-tuning.
Outcome: The proposed framework improves multi-domain fostering performance by 29.77% compared to uniform weights.
Teaching Language Models to Self-Improve by Learning from Language Feedback (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) generate content that can be untruthful or harmful.
Approach: They propose a method that leverages model feedback for alignment . they use a base language model to generate initial responses, critiqued and refined .
Outcome: The proposed method outperforms strong baselines across diverse tasks and model sizes.
Look Before You Leap: A Lookahead Reasoning Quality Gate for Speculative Decoding (2026.eacl-long)

Copied to clipboard

Challenge: Unlike token-level likelihood search, which is myopic and often rewards verbosity, our approach works at an intermediate granularity.
Approach: They propose a lookahead quality gate for speculative decoding that accepts the longest reliable prefix of each k-token lookaheaded draft.
Outcome: The proposed method improves accuracy over baselines while achieving 2.6-7.9 faster generation on math and science benchmarks.
ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning (2026.acl-long)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) is a powerful framework for knowledge-intensive tasks, but its effectiveness in long-context scenarios is often bottlenecked by the retriever’s inability to distinguish sparse yet crucial evidence.
Approach: They propose a framework that fine-tunes the retriever for Answer Alignment by identifying high-quality positive chunks by evaluating their sufficiency to generate the correct answer.
Outcome: The proposed framework improves 14.5% over the base model and maintains strong efficiency for long-context RAG.
Does Meta-learning Help mBERT for Few-shot Question Generation in a Cross-lingual Transfer Setting for Indic Languages? (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to few-shot Question Generation (QG) are limited and require manual annotation.
Approach: They propose to use multilingual BERT to perform few-shot question generation with cross-lingual transfer.
Outcome: The proposed model improves in few-shot QG and human evaluation confirms it.
Translation and Fusion Improves Cross-lingual Information Extraction (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown significant progress in information extraction tasks due to lack of labeled data for fine-tuning and unlabeled text for pre-training.
Approach: They propose a framework in which large language models are fine-tuned to use English translations of low-resource language data.
Outcome: The proposed model improves cross-lingual transfer over the base model on 12 multilingual IE datasets spanning 50 languages.
Contrastive Learning for Prompt-based Few-shot Language Learners (2022.naacl-main)

Copied to clipboard

Challenge: a recent study has shown that GPT-3 fine-tuning models with limited examples is effective . a contrastive learning framework clusters inputs from the same class under different augmented “views” and repels those from different classes.
Approach: They propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" they combine a contrastive loss with the standard masked language modeling loss in prompt-based few-shot learners .
Outcome: The proposed framework improves on the state-of-the-art methods in a diverse set of 15 language tasks.
Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset (2020.acl-main)

Copied to clipboard

Challenge: Medical professionals often query over clinical notes to find information that can support their decision making.
Approach: They propose to use expert-annotated question templates and existing i2b2 annotations to create emrQA, the first large-scale dataset for question answering based on clinical notes.
Outcome: The proposed system can answer clinical questions without using domain knowledge.
Counterfactual Inference for Text Classification Debiasing (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to capture unintended dataset biases are expensive and require elaborate balancing strategies.
Approach: They propose a model-agnostic text classification debiasing framework which can effectively avoid employing data manipulations or designing balancing mechanisms.
Outcome: The proposed framework can effectively avoid data manipulations or designing balancing mechanisms to capture unintended dataset biases.
Experience-driven Multi-turn Reinforcement Learning for GUI Agents (2026.acl-long)

Copied to clipboard

Challenge: GUI agents have demonstrated remarkable progress in automating complex user interface interactions . training such agents for long-horizon tasks remains challenging due to limited rewards and prohibitive costs.
Approach: They propose a method that leverages expert trajectories as environment experiences for on-policy multi-turn training.
Outcome: The proposed method achieves significant gains over the base model with 1K public trajectories as RL experiences . it achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o .
Can Explanations Be Useful for Calibrating Black Box Models? (2022.acl-long)

Copied to clipboard

Challenge: Existing models are often used as black boxes to adapt to new domains, but there is no single recipe for making them work.
Approach: They propose to use black box models to improve their performance on new domains by leveraging explanations of their behavior.
Outcome: The proposed method improves model generalization performance on two tasks using explanations.
Efficient End-to-End Visual Document Understanding with Rationale Distillation (2024.naacl-long)

Copied to clipboard

Challenge: Pre-processing tools such as optical character recognition (OCR) can map document image inputs to textual tokens, then large language models (LLMs) can reason over text.
Approach: They propose a method that integrates outputs of OCR tools and larger multimodal models as intermediate "rationales" a student model is trained to predict rationales and answers based on visual documents .
Outcome: The proposed model outperforms the base model on three visual document understanding benchmarks with only 1% higher computational cost.
EC-RAFT: Automated Generation of Clinical Trial Eligibility Criteria through Retrieval-Augmented Fine-Tuning (2025.findings-acl)

Copied to clipboard

Challenge: Eligibility criteria (EC) are critical components of clinical trial design, specifying parameters for participant inclusion and exclusion.
Approach: They propose a method that utilizes Retrieval-Augmented Fine-Tuning to generate structured and cohesive EC directly from clinical trial titles and descriptions.
Outcome: The proposed method outperforms Llama-3.1-8B-Instruct and Llm-as-a-Judge models in BERTScore and EC score.
Beyond Pedagogical Principles: Multi-Horizon Preference Optimization for Efficient Socratic Tutoring (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for developing LLMs are constrained by static data or sparse reward signals in online settings.
Approach: They propose a framework that iteratively refines tutor agents using a multi-horizon reward function within a dynamic teacher-student simulation environment.
Outcome: The proposed framework improves model performance and balances principles and effectiveness compared to baselines.
SimpleOCR: Rendering Visual Questions to Teach MLLMs to Read (2026.findings-acl)

Copied to clipboard

Challenge: MLLMs lack visual grounding mechanism to read text embedded in images, or rely on parametric shortcuts . despite strong OCR capabilities, models suffer performance degradation of 12.7% in the VQ setting .
Approach: They propose a plug-and-play training strategy that invalidates shortcuts in text prompts . they propose 'vq' setting where text queries are rendered directly onto images .
Outcome: The proposed training strategy surpasses the base model by 5.4% and GRPO based on original images by 2.7% on four representative OOD benchmarks.
Fingerprinting LLMs via Prompt Injection (2026.acl-long)

Copied to clipboard

Challenge: Existing provenance detection methods for large language models are infeasible for already published models and compare outputs using hand-crafted or random prompts.
Approach: They propose a detection framework that constructs fingerprints by exploiting LLMs’ inherent vulnerability to prompt injection.
Outcome: The proposed framework achieves high true positive rates while keeping false positive rates near zero.
DistillCSE: Distilled Contrastive Learning for Sentence Embeddings (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to sentence embeddings are based on contrastive learning (CL) .
Approach: They propose a framework which performs contrastive learning under the self-training paradigm with knowledge distillation and propose 'Group-P shuffling strategy' and averaging logits from multiple teacher components.
Outcome: The proposed framework outperforms many strong baseline methods and yields a new state-of-the-art performance.
RQT: Hierarchical Residual Quantization for Multi-Model Compression (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for decomposing fine-tuned LLMs are sensitive to the magnitude of delta values.
Approach: They propose a hierarchical quantization framework that shares low-bit integer weights across similar models.
Outcome: The proposed framework achieves an average accuracy degradation of approximately 3% on fine-tuned models across mathematics, coding, chatbot, and Chinese LLMs.
Rethinking Attribute Representation and Injection for Sentiment Classification (D19-1)

Copied to clipboard

Challenge: Existing models that use text attributes to improve sentiment classification use text as a categorical feature.
Approach: They propose to represent attributes as chunk-wise importance weight matrices and consider four locations to inject attributes.
Outcome: The proposed method outperforms the state-of-the-art and outperformed previous models.
Speechworthy Instruction-tuned Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Current instruction tuned language models are trained on textual preference data and therefore not aligned to speech domain.
Approach: They propose to use radio-industry best practices to prompt and learn speech-based preference data to improve speech-suitability of popular instruction tuned language models.
Outcome: The proposed methods achieve the best win rates in head-to-head comparisons, resulting in preferred or tied to the base model in 76.2% of comparisons on average.
Exploring Design Choices for Building Language-Specific LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior work focused on building multilingual models that cover a broad spectrum of languages.
Approach: They conduct systematic experiments on how design choices impact the adapted LLM, both in terms of efficiency and end task performance.
Outcome: The proposed model performs better on English-centric models than multilingual models despite poor performance on low-resource languages.
An Empirical Comparison on Imitation Learning and Reinforcement Learning for Paraphrase Generation (D19-1)

Copied to clipboard

Challenge: Existing methods to generate paraphrases are not trivial and often fail in practice.
Approach: They propose to use imitation learning to boost the performance of generating paraphrases by using a pointer-generator model.
Outcome: The proposed model outperforms the state-of-the-art methods on the benchmark datasets.
UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision (2026.acl-long)

Copied to clipboard

Challenge: Unified Multimodal Models have achieved remarkable success in cross-modal comprehension, but a gap persists in their ability to translate internal knowledge into faithful and controllable synthesis.
Approach: They propose a self-improvement framework that partitions a single UMM into three collaborative roles: Proposer, Solver, and Judge.
Outcome: The proposed framework improves on TIIF, DPG, CompBench and UniCycle benchmarks.
Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) increase test-time computation, often in the form of chain-of-thought (CoT) however, reasoning traces can become unnecessarily long, increasing computation costs without improving accuracy and sometimes even degrading performance.
Approach: They propose a multi-stage efficient reasoning method that combines supervised fine-tuning with reinforcement learning using an adaptive length penalty.
Outcome: The proposed method reduces response length by an average of 28% for 8B models and 40% for 32B models while incurring only minor performance drops of 1.6 and 2.5 points, respectively.
Nudging: Inference-time Alignment of LLMs via Guided Decoding (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) require alignment to effectively and safely follow user instructions.
Approach: They propose a simple, training-free algorithm that aligns any base model at inference time using a small aligned model.
Outcome: The proposed algorithm outperforms large aligned models on open-instruction tasks without training.
You Only Need One Single Token to Refine Safety Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Excessive safety can lead to over-refusal, where models reject harmful-looking yet benign queries, severely limiting utility.
Approach: They propose a lightweight training-based approach that reshapes the distributions of harmful and benign samples within the model’s decision space by using a single-token prefix.
Outcome: The proposed approach can distinguish between harmful and benign samples while keeping the model frozen.
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress? (2024.emnlp-main)

Copied to clipboard

Challenge: Several studies claim that domain-adaptive pretraining improves performance on downstream medical tasks.
Approach: They compare medical LLMs and VLMs against their corresponding base models . they find that medical Lms outperform their base models in 12.1% of cases .
Outcome: The proposed models outperform their base models on medical questions and tasks in 12.1% of cases and reach a tie in 49.8% of cases.
CaPE: Contrastive Parameter Ensembling for Reducing Hallucination in Abstractive Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing work suggests that the degree of hallucination depends on factual errors in training data.
Approach: They propose a method to use training data to reduce hallucination by ensembling parameter variations in training data.
Outcome: The proposed method improves on XSUM and CNN/DM datasets on human evaluations and factual metrics.
Large Language Models Can Not Perform Well in Understanding and Manipulating Natural Language at Both Character and Word Levels? (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) still exhibit significant deficiencies in basic language understanding and manipulation.
Approach: They propose a bilingual benchmark to assess the performance of Large language models . they use a set of 15 simple text editing tasks to examine their capabilities .
Outcome: The proposed benchmark aims to assess the performance of Large language models in basic language tasks.
StyleBART: Decorate Pretrained Model with Style Adapters for Unsupervised Stylistic Headline Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on unsupervised headline generation focus on a standard dataset and mono-style corpora.
Approach: They propose an unsupervised approach for stylistic headline generation using a pretrained BART model decorated with adapters responsible for different styles.
Outcome: The proposed method separates the task of style learning and headline generation, allowing for the generation of diverse headlines with diverse styles.
Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models (2025.acl-long)

Copied to clipboard

Challenge: Parameter-efficient fine-tuning (PEFT) is a common method for fine- tuning large language models . however, once updated, PEFT modules suffer performance degradation on newer versions .
Approach: They propose a method that enhances the PEFT module by focusing on the task-specific pattern while reducing its dependence on certain knowledge in the base model.
Outcome: Experiments show that PEFT modules can maintain performance on updated models without re-tuning . the proposed approach can be used in real-world applications with large model sizes .
Learning to Translate by Translating: Stabilizing the Dual Loop via Semantic-Aware Self-Evolution (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been successful in machine translation, but lack of high-quality parallel corpora and cost constrain scalability.
Approach: They propose an LLM-driven dual-learning framework that enables autonomous translation . they employ a robust semantic-aware reward function that balances adequacy with reconstruction fidelity .
Outcome: The proposed model outperforms larger models on benchmarks and achieves parity with state-of-the-art supervised baselines on mainstream benchmarks.
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings (2025.emnlp-main)

Copied to clipboard

Challenge: Reasoning-capable large language models (LLMs) have driven a major shift in artificial intelligence . these models generate long CoTs, capturing reasoning behaviors such as self-reflection, self-correction, and hypothesis testing.
Approach: They propose a sample-efficient, two-stage training strategy to build reasoning LLMs . they "warm up" a model by distilling Long CoTs from a toy domain to acquire general reasoning skills .
Outcome: The proposed training strategy outperforms existing models on a range of tasks.
MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap.
Approach: They propose a benchmark that links MIMIC-IV EHRs to a unified knowledge base built from UMLS and other biomedical vocabularies.
Outcome: The proposed model improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors.
Controllable Style Arithmetic with Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for linguistic style control lack fine-grained control, require extensive computation, or introduce significant latency.
Approach: They propose a parameter-space approach that extracts style-specific representations by analyzing parameter differences between models trained on contrasting styles and incorporates them into a model with precise control over style intensity.
Outcome: The proposed approach achieves three key capabilities while achieving optimal computational efficiency.
End-to-End Bias Mitigation by Modelling Biases in Corpora (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have shown that strong natural language understanding models are prone to relying on unwanted dataset biases without learning the underlying task.
Approach: They propose two learning strategies to train neural models that are more robust to dataset biases and transfer better to out-of-domain datasets.
Outcome: The proposed methods improve robustness in all settings and transfer better to out-of-domain datasets.
Enhancing Training Data Attribution for Large Language Models with Fitting Error Consideration (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are difficult to interpret due to their black-box nature and randomness.
Approach: They propose a new method which enhances influence functions by addressing fitting errors by eliminating knowledge bias present in the base model before fine-tuning.
Outcome: The proposed method outperforms existing methods and achieves an average AUC of 91.64%.
Incorporating Lexical and Syntactic Knowledge for Unsupervised Cross-Lingual Transfer (2024.lrec-main)

Copied to clipboard

Challenge: Unsupervised cross-lingual transfer is a process of transferring knowledge between languages without explicit supervision.
Approach: They propose a framework that combines lexical and syntactic knowledge to enhance learning . they use a code-switching technique to implicitly teach lexica and a syntaktic-based graph attention network to help encode syntakic structure.
Outcome: The proposed framework outperforms baselines of zero-shot cross-lingual transfer with 1.0 3.7 points on text classification, named entity recognition, and semantic parsing tasks.
Enhancing Document-level Event Argument Extraction with Contextual Clues and Role Relevance (2023.findings-acl)

Copied to clipboard

Challenge: Document-level event argument extraction is a challenging task for cross-sentence inference . previous work focused on document-level EAE, but recent work focused more on documentlevel .
Approach: They propose a document-level event argument extraction model that captures contextual clues and latent role information.
Outcome: The proposed model outperforms existing methods on two public datasets with 1.13 F1 and 2.64 F1 improvements on RAMS and WikiEvents respectively.
Token-level Preference Self-Alignment Optimization for Multi-style Outline Controllable Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing attempts to outline generation are limited by response pair requirements and substantial computation costs.
Approach: They propose a token-level preference self-alignment optimization for outline controllable generation that extends the Bradley-Terry model from pair-wise to list-wise comparison.
Outcome: The proposed method outperforms existing methods by 19.28% in performance while requiring only 56.25% training time.
BOSE: A Systematic Evaluation Method Optimized for Base Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods for large language models (LLMs) are inadequate to provide solid conclusions for key experiments such as data ablation and scaling law.
Approach: They propose a method specifically designed to optimize the evaluation of base models by incorporating two innovations: In-Context Light-instruction Prompt and Blank-ppl for multi-choice tasks with candidate options.
Outcome: The proposed method significantly improves stability and consistency of evaluations during pre-training and consistency between base and instruct models.
Let’s Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM’s Math Capability (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work has focused on improving the mathematical reasoning capabilities of Large Language Models (LLMs).
Approach: They propose an end-to-end framework to integrate FL into NL math reasoning . they propose a problem alignment method that reformulates QA and existence problems .
Outcome: The proposed framework achieves 89.80% and 84.34% accuracy rates on the MATH-500 and the AMC benchmarks.
LexMatcher: Dictionary-centric Data Curation for LLM-based Machine Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: emergence of large language models (LLMs) has brought about new opportunities for machine translation.
Approach: They propose a method for data curation that supplements the infrequent senses of polysemous words.
Outcome: The proposed method outperforms established baselines on the WMT2022 test sets and is applicable to other pre-trained models.
LightVLP: A Lightweight Vision-Language Pre-training via Gated Interactive Masked AutoEncoders (2024.lrec-main)

Copied to clipboard

Challenge: Existing vision-language pre-training models use multi-modal encoders to encode image and text, causing noisy training corpora.
Approach: They propose a vision-language pre-training framework with two autoencoders for efficient training . they propose masked tokens and a gated interaction mechanism to cope with noise .
Outcome: The proposed model achieves 2.2% R@1 gains on COCO Text Retrieval and 1.1% on refCOCO+ on six datasets.
Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extraction (2025.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to structured NLP tasks use autoregressive models trained on pairs of unstructured input text and structured output targets.
Approach: They propose a model that combines constrained and unconstrained decoding in two phases to achieve two weak predictions.
Outcome: The proposed model outperforms previous approaches both in and out of distribution, addressing several common errors identified in those approaches.
RA-LoRA: Rank-Adaptive Parameter-Efficient Fine-Tuning for Accurate 2-bit Quantized Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) with their extensive parameters and high memory demands are challenging to fine-tune for specific applications with limited resources.
Approach: They propose a method that dynamically adjusts the adapter’s rank using rank-subspace analysis, optimizing performance with fewer parameters.
Outcome: The proposed method improves model accuracy with minimal parameter changes and demonstrates the importance of rank dynamics in optimizing quantized LLMs.
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment (2025.acl-long)

Copied to clipboard

Challenge: Large language models such as GPT-4 have limited their deployment in clinical settings . a novel framework for adapting SLMs into high-performing clinical models is needed .
Approach: They propose a framework for adapting large language models into high-performing clinical models . they pre-instruct experts on relevant medical and clinical corpora and model merging .
Outcome: The proposed framework outperforms the existing model on the CLUE+ benchmark on medical entities and radiology reports.
PhonoThink: Improving Large Language Models’ Reasoning on Chinese Phonological Ambiguities (2025.emnlp-main)

Copied to clipboard

Challenge: Effectively resolving phonological ambiguities is crucial for robust natural language processing, as these ambiguity are pervasive in tasks ranging from speech-to-text, spelling correction, to offensive language detection.
Approach: They propose a framework to enhance LLMs’ phonological capability through a multiple-stage training approach.
Outcome: The proposed framework enables the base model to achieve comparable performance to a much larger model.
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention (2025.emnlp-main)

Copied to clipboard

Challenge: Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation.
Approach: They propose a method that integrates speculative draft generation directly within the target model using multi-stream attention.
Outcome: The proposed method improves acceptance but also latency and speculation latency, limiting overall speedup.
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing RL methods suffer from reliability bottlenecks due to reward sparsity and intractable computations . d-TreeRPO provides fine-grained and verifiable step-wise reward signals .
Approach: They propose a reliable reinforcement learning framework for diffusion large language models that leverages tree-structured rollouts and bottom-up advantage computation based on verifiable outcome rewards.
Outcome: The proposed framework outperforms baseline models and achieves significant improvements across reasoning benchmarks.
Mitigating Catastrophic Forgetting in Language Transfer via Model Merging (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models have shown remarkable capabilities, particularly in English, but for less prevalent languages, performance can be significantly lower, making additional adaptation paramount.
Approach: They propose a new adaptation method based on iteratively merging multiple models fine-tuned on a subset of available training data that reduces forgetting while maintaining learning on the target domain.
Outcome: The proposed method outperforms LLAMA-3-8B-based models in German and German while maintaining learning on the target domain.
TESS 2: A Large-Scale Generalist Diffusion Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing instruction-following diffusion models are predominantly trained using an autoregressive paradigm.
Approach: They propose a general instruction-following diffusion language model that outperforms contemporary instruction-tuned diffusion models and matches and sometimes exceeds strong autoregressive (AR) models.
Outcome: The proposed model outperforms and sometimes exceeds existing autoregressive (AR) models on a number of tasks.
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies address safety-constrained online and offline preferences optimizations, but offline methods perform poorly in adaptively balancing safety and helpfulness.
Approach: They propose a mixture of experts framework for safety-helpfulness dual Preference Optimization . they combine a single-preference enhanced direct preference optimization approach with a dynamic routing mechanism .
Outcome: The proposed framework outperforms state-of-the-art methods in safety and helpfulness.
Beyond Modality Collapse: Taming Guided Modality Entropy for Omni-modal Emotion Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: EmoOmni is a data paradigm for omni-modal large language models that can be used for emotion reasoning.
Approach: They propose a data paradigm that interleaves guided tokens into reasoning traces to enforce structured evidence extraction.
Outcome: The proposed paradigm over-relys on a dominant modality while neglecting complementary cues.
The Emperor’s New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFT (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have yielded impressive gains on mathematical reasoning benchmarks via supervised fine-tuning (SFT).
Approach: They investigate the mechanisms behind SFT improvements in small-scale large language models by examining four key questions: (1) Are performance gains primarily due to format alignment rather than reasoning? (2) Can high-quality supervision encourage genuine reasoning? (4) Are format alignment gains consistent across model sizes and architectures?
Outcome: The proposed models outperform the proprietary models on OlympiadBench and Omni-Math, but lack the brittleness of the models under perturbations to test their reasoning abilities.
DEM: Distribution Edited Model for Training with Mixed Data Distributions (2024.emnlp-main)

Copied to clipboard

Challenge: Recent fine-tuning approaches for large language models require supervised finetun on diverse datasets and follow different distributions.
Approach: They propose a distribution edited model that integrates models individually trained on each data source with the base model using basic element-wise vector operations.
Outcome: The proposed model outperforms baseline models on a variety of benchmarks and is cheaper than standard data mixing methods.
zFLoRA: Zero-Latency Fused Low-Rank Adapters (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed with task-specific adapters catering to multiple downstream applications.
Approach: They propose a low-latency fused low-rank adapter that introduces zero latency overhead on top of the base model.
Outcome: The proposed adapter reduces the inference time of the model by 2.5x . the proposed adapters are tested on 18 different tasks on different platforms .
PAD: A Robustness Enhancement Ensemble Method via Promoting Attention Diversity (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to enhance robustness of deep neural networks focus on perturbation . weak robustness is a problem for many types of adversarial attacks, authors say .
Approach: They propose a lightweight framework for enhancing robustness by perturbing parameters of a model and diversifying adversarial example distributions among different models.
Outcome: The proposed method can improve robustness against adversarial attacks while maintaining accuracy on clean data.
SecDecoding: Steerable Decoding for Safer LLM Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing decoding-time defense methods suffer from limited generalization, high computational overhead, or significant utility degradation.
Approach: They propose a decoding-time defense framework that leverages a pair of small contrastive models to estimate token-level safety signals by measuring divergence in their output distributions.
Outcome: The proposed framework achieves near-zero attack success rates against a wide spectrum of advanced jailbreak attacks while maintaining the model’s helpfulness with minimal degradation.
Adversarial Preference Learning for Robust LLM Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Modern language models rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors, but they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation; (2) the vast diversity of potential adversarials; and (3) the risk of feedback bias and reward hacking.
Approach: They propose an iterative adversarial training method that incorporates three key innovations to address these challenges.
Outcome: Experiments on Mistral-7B-Instruct-v0.3 show that the proposed method significantly enhances robustness and reduces harmful outputs from 5.88% to 0.43%.
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning (2025.acl-long)

Copied to clipboard

Challenge: Representation Fine-tuning (ReFT) is a proposed method for improving parameter efficiency . however, it yields suboptimal performance, as fixed-position representations have uncertain impact on outputs .
Approach: They propose a method that fine-tunes critical representations in a low-rank linear subspace while freezing the base model.
Outcome: The proposed method improves accuracy of LLaMA-2-7B and ReFT by 18.2 and 3.8 on GSM8K.
Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings (2025.emnlp-main)

Copied to clipboard

Challenge: Modern document retrieval embedding methods typically encode passages (chunks) from documents independently, often overlooking contextual information from the rest of the document.
Approach: They propose a benchmark to evaluate retrieval models' ability to leverage document-wide context.
Outcome: The proposed method significantly improves retrieval quality on ConTEB without sacrificing base model performance.
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for long-horizon agents introduce the external memory module and look up the relevant information from the stored memory, which prevents the model from proactively managing its memory content and aligning with the agent’s overarching task objectives.
Approach: They propose an algorithm which enables agents to autonomously manage their memory during interaction with environment and selectively retain crucial information.
Outcome: Extensive experiments show that the proposed algorithm achieves absolute F1 score gains of 25.98 over the base model and 7.1 over the previous SOTA baseline while preserving task performance.
LoRE-Merging: Exploring Low-Rank Estimation For Large Language Model Merging (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework for model merging is proposed without additional training . task vectors from fine-tuned models exhibit a limited number of dominant singular values .
Approach: They propose a framework for model merging based on low-rank estimation of task vectors without access to the base model.
Outcome: The proposed framework improves models without additional training without additional inputs.
FiRST: Finetuning Router-Selective Transformers for Input-Adaptive Latency Reduction (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to improve latency via skipping layers have limitations . fiRST is a model-agnostic framework that reduces inference latency while maintaining quality .
Approach: They propose a model-agnostic framework that skips transformer layers during decoding . it is fully compatible with KV caching, enabling faster decoding while maintaining quality .
Outcome: a new framework reduces inference latency by using layer-specific routers to skip transformer layers during decoding.
Efficient Ensemble for Fine-tuning Language Models on Multiple Datasets (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning language models are efficient when adapting to a single dataset.
Approach: They propose to use an ensemble method for fine-tuning a language model to multiple datasets instead of a single adapter per task.
Outcome: The proposed method improves performance on multiple datasets while preserving low-rank adaptation properties.
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning (2025.acl-long)

Copied to clipboard

Challenge: Results show that supervised fine-tuning and preference finetunation are the most efficient approaches for large language models.
Approach: They propose to use Supervised Finetuning and Preference Finetunes to optimize training data budgets for Large Language Models.
Outcome: The proposed approach improves performance on math tasks by 15% on the most expensive model, 1,000 examples.
HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches focus on retrieval augmentation and focus on the quality of the output . Existing methods focus on generating a highly specific declarative statement ignoring the underlying reasoning process behind ideation.
Approach: They propose a large language model that generates evidence-based hypotheses using literature-guided reasoning and a multi-task setting.
Outcome: The proposed model outperforms the base model and generates evidence-grounded hypotheses with high feasibility and impact as judged by human experts.
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening (2025.emnlp-main)

Copied to clipboard

Challenge: Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities.
Approach: They propose a method for explicitly up-weighting rare but correct solutions to overcome rank bias in group relative policy optimization (GRPO) .
Outcome: The proposed method mitigates rank bias and improves pass@N across a large range of N in both synthetic and real theorem proving settings.
Breaking Token Into Concepts: Exploring Extreme Compression in Token Representation Via Compositional Shared Semantics (2025.findings-emnlp)

Copied to clipboard

Challenge: Standard language models employ unique, monolithic embeddings for each token, limiting their ability to capture multifaceted meanings.
Approach: They propose a compositional structure that accumulates diverse semantic facets for tokens . they apply this representational scheme to standard transformer architectures and a biomedical domain benchmark .
Outcome: The proposed representational scheme achieves extreme compression in embedding parameters while maintaining >95% task performance relative to the base model.
Will it Merge? On The Causes of Model Mergeability (2026.findings-acl)

Copied to clipboard

Challenge: Model merging has emerged as a promising technique for combining fine-tuned models into a single expert model without retraining.
Approach: They propose a model merging technique that preserves weak model knowledge . they define mergeability as a property of model updates that captures how well they retain trained knowledge when merged with other model updates.
Outcome: The proposed method preserves weak knowledge in the base model.
Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut Decoding (2026.acl-long)

Copied to clipboard

Challenge: Existing methods that prune or employ early stopping to reduce latency often compromise reasoning reliability.
Approach: They propose a shortcut decoding framework that integrates probes over internal hidden states with step-level entropy to detect convergence of reasoning during generation and adaptively selects between a fast-exit path and a stability-verified path to remove redundant steps while preserving answer correctness.
Outcome: The proposed framework reduces token usage by approximately 35% and maintains accuracy comparable to full CoT decoding.
OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization (2025.acl-long)

Copied to clipboard

Challenge: Existing studies focus on building text-only agents in synthetic environments where the reward signals are clearly defined.
Approach: They propose a multimodal web agent that can autonomously conduct real-world exploration and improve itself after each iteration.
Outcome: The proposed agent improves itself after each iteration, demonstrating strong performance across multiple test sets.
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety defenses typically intervene internally within the generative model, but suffer from severe concept entanglement, leading to degradation of benign generation quality.
Approach: They propose a structurally isolated safety module that performs external, interpretable rectification without modifying the base model.
Outcome: The proposed module performs external, interpretable rectification without modifying the base model.
Tailored Primitive Initialization is the Secret Key to Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Reinforcement learning (RL) has emerged as a powerful paradigm for improving the reasoning capabilities of large language models.
Approach: They propose a pipeline that automatically discovers thinking token patterns with reasoning primitives and curates SFT datasets to prepare LLMs for RL.
Outcome: The proposed pipeline outperforms baseline methods on mathematical and logical reasoning benchmarks on RL tasks.
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm .
Approach: They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks.
Outcome: The proposed model outperforms the current state-of-the-art in label detection and explanation generation.
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists (2025.emnlp-main)

Copied to clipboard

Challenge: AutoSDT-5K is the only automatically collected and the largest open dataset for data-driven scientific discovery.
Approach: They propose an automatic pipeline that collects high-quality coding tasks in real-world data-driven discovery workflows.
Outcome: The proposed pipeline synthesizes accurate tasks and tasks from a dataset of 5,404 tasks covering four scientific disciplines and 756 Python packages.
Grouped Adaptive Weight Sharing (GAWS): An Inference-Efficient Adaptation Method for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) is a new approach to fine-tuning large language models . adapters are lightweight, task specific modules that can be used for adapters in latency-sensitive settings.
Approach: They propose a low-rank adapter with a weight sharing mechanism that reduces latency by 40% . they analyze LoRA adapters on GPUs and identify segmented function calls as the primary source of latency.
Outcome: The proposed adapter reduces latency to about 40% of the gap between the unmerged LoRA and the base model while maintaining parameter efficiency and comparable accuracy.
Auto-Stega: An Agent-Driven System for Lifelong Strategy Evolution in LLM-Based Text Steganography (2026.findings-acl)

Copied to clipboard

Challenge: prevailing methods rely on hand-crafted or pre-specified strategies and struggle to balance efficiency, imperceptibility, and security, particularly at high embedding rates.
Approach: They propose an agent-driven self-evolving framework that is the first to realize self-changing steganographic strategies by automatically discovering, composing, and adapting strategies at inference time.
Outcome: The proposed framework achieves 42.2% perplexity and 1.6% anti-steganalysis performance over SOTA methods at high embedding rates.
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning . however, the recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies.
Approach: They propose a replay strategy with dynamic objective reweighting for general knowledge preservation using short-horizon signals of convergence and instability.
Outcome: The proposed method preserves general capabilities and improves reasoning . it can be applied to existing RLVR pipelines without training additional models or tuning .
Better LLM Reasoning via Dual-Play (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR) however, external supervision remains a bottleneck for tasks and domains for which supervised data are scarce or non-existent.
Approach: They propose a novel dual-play framework that adversarially trains two models initialized from the same base model.
Outcome: The proposed framework improves the math reasoning performance of large language models.
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Recent Large Audio Language Models (LALMs) have shown strong capabilities in audio understanding, yet their reasoning remains vulnerable to perceptual errors.
Approach: They propose a large-scale dataset for **Perception-Aware Question Answering** that uses a hierarchical decoupling strategy to separate speech from environmental sounds and distinguishes among multiple speakers.
Outcome: The proposed model improves on MMAU-mini, MMAR, and PAQA while maintaining comparable performance on multiple benchmarks.
Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting.
Approach: They propose a feature-level mechanistic analysis methodology to probe RL generalization using a controlled experimental setup.
Outcome: The proposed method identifies a compact, task-agnostic set of features that directly mediate generalization across diverse tasks.
Figure It Out: Improve the Frontier of Reasoning with Executable Visual States (2026.acl-long)

Copied to clipboard

Challenge: Recent reasoning models fail to capture structural constraints in complex settings.
Approach: They propose a visual-based reasoning system that integrates executable visual construction into multi-turn reasoning via end-to-end reinforcement learning.
Outcome: The proposed model outperforms strong text-only chain-of-thought models on seven mathematical benchmarks and improves by 13.12% on AIME 2025 and 11.00% on BeyondAIME.
Why Can Distillation Work with Limited Resources? A Systematic Study (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have driven reasoning performance . low-resource distillation can boost models' performance, but a framework is missing .
Approach: They conduct a controlled experiment to find out why low-resource distillation can boost model performance . they find that distillation enhances the presence of advanced cognitive behaviors .
Outcome: The proposed model shows more flexible reasoning, the authors show . they show that distillation enhances the presence of advanced cognitive behaviors .
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks lack systematically paired instances across modalities, making it difficult to compare genuine arithmetic limits . a model that computes 4736 may fail on a nearby instance like 8967, despite a well-tuned internal router.
Approach: They propose a controlled multimodal multiplication benchmark that factorially varies digit length, digit sparsity, representation, and modality with paired instances from a reproducible generator.
Outcome: The proposed model can perceive numerical content across modalities but fails to perform exact multi-digit multiplication when presented as numerals, number words, images, or in audio form.
SCALE: Upscaled Continual Learning of Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Recent discussions suggest that further progress will come from scaling the right structure, not merely parameters or data, while preserving acquired knowledge.
Approach: They propose a width upscaling architecture that inserts lightweight expansions into linear modules while freezing all pre-trained parameters.
Outcome: The proposed architecture reduces severe forgetting while learning new knowledge on a controlled synthetic biography benchmark.
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for text-to-image alignment evaluation rely on coarse-grained metrics or static Question Answering pipelines that lack fine-grounded interpretability and struggle to reflect human preferences.
Approach: They propose a reinforcement-guided visual reasoning framework for element-level text-to-image alignment evaluation.
Outcome: The proposed framework achieves state-of-the-art results on four benchmarks and surpasses the strong proprietary Gemini 3 Pro and Training-based baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations